Security Architecture
Health systems are attacked because health data is valuable and health organisations pay ransoms. Ransomware against hospitals is now a routine operational risk with documented effects on patient outcomes, not a hypothetical one, and the architecture has to assume compromise rather than prevent it absolutely.
This page covers the technical security architecture. The management system around it is ISO 27001; health-specific controls are ISO 27799; the legal layer is GDPR and its national equivalents.
Defence in depth
No single control is sufficient. The layers, and the health-specific note for each:
| Layer | Controls | Health-specific note |
|---|---|---|
| Governance | Policy, risk register, ownership, training | Clinical staff are the largest user population and the least available for training |
| Identity | MFA, OIDC, privileged access management | Shared workstations and shift working break naive MFA designs |
| Network | Segmentation, no flat networks, egress control | Medical devices are frequently unpatchable and must be isolated, not fixed |
| Application | Secure development, dependency scanning, code review | Long-lived clinical systems accumulate unsupported dependencies |
| Data | Encryption, minimisation, masking, retention | Retention periods are decades; key management must outlive staff |
| Detection | Logging, SIEM, alerting, anomaly detection | Access-pattern monitoring is the main control against insider misuse |
| Response | Incident plan, rehearsed, with clinical continuity | The plan must specify how care continues while systems are down |
| Recovery | Backups, tested restores, immutable copies | Restore time is a clinical safety parameter |
Encryption
In transit
TLS 1.2 minimum, 1.3 preferred, everywhere — including inside the data centre, which is what zero trust means in practice. Mutual TLS for system-to-system exchange where client certificates are manageable.
Legacy health protocols need attention: HL7 v2 over MLLP has no native encryption and is usually tunnelled or run over a protected network. DICOM's classic DIMSE services are similar — see DICOM.
At rest
Full-disk or volume encryption is table stakes and protects against physical theft only. It does nothing against a compromised application or a stolen credential, which is how data actually leaves.
Application-level or column-level encryption for the most sensitive fields provides real defence in depth, at the cost of losing the ability to query those fields. Decide per field, deliberately.
Backups must be encrypted with keys held separately from the production key store — otherwise a compromise that reaches the key store reaches the backups too.
Key management
The part that is usually improvised and should not be:
- Keys in a dedicated key management service or HSM, never in application configuration or source control
- Documented rotation schedule, with rotation actually performed and rehearsed
- Separation of duties between those who administer keys and those who administer data
- Key escrow and recovery — health data has retention periods measured in decades; the person who generated the key will have left. A recovery procedure that has never been tested is not a recovery procedure.
- Key destruction as a deliberate, logged step at end of retention
Secrets management
Database passwords, API credentials, signing keys. Requirements: a central store (Vault/OpenBao or a cloud equivalent), short-lived dynamic credentials where possible, no secrets in source control, automated scanning to catch the ones that get committed anyway, and rotation on staff departure.
PKI and certificates
Multi-organisation exchange usually needs a certificate authority: organisation identity, mutual TLS, document and message signing.
The recurring failure is expiry. A certificate expiring at 02:00 takes down the national exchange, and the person who issued it left two years ago. Mitigations: automated issuance and renewal (ACME, cert-manager, step-ca), monitoring with alerts at 90, 30 and 7 days, a complete inventory of every certificate including those inside appliances, and a documented emergency issuance procedure.
Threat modelling
Do it per component, at design time. STRIDE is a serviceable framework, and the health-specific threats worth naming explicitly:
| Threat | Realistic scenario |
|---|---|
| Insider curiosity | Staff looking up a neighbour, a colleague, or a public figure. The most common privacy breach by a wide margin, and it is detected by audit analysis, not by prevention. |
| Credential compromise | Phishing against clinical staff; shared workstation credentials |
| Ransomware | Encryption of clinical systems and their backups, with a demand |
| Supply chain | A compromised vendor with remote support access into the clinical network |
| Medical device compromise | Unpatchable equipment as a foothold |
| Bulk exfiltration | A bulk export endpoint with broad scopes |
| Re-identification | "De-identified" data linked against other sources |
| Availability attack | Denial of service during a clinical emergency |
| Data integrity | Silent alteration of results — rarer, and the worst case clinically |
Note that several of these are not prevented by perimeter controls at all. Insider misuse and bulk exfiltration are detected in the audit log or not at all, which is why audit is an architecture requirement rather than a compliance checkbox.
Logging and audit
Two distinct things, often conflated:
- System logs — for operating and debugging the software. Must not contain personal health data.
- Audit records — who accessed which patient's data, when, for what purpose,
and whether it was permitted. FHIR
AuditEvent/ IHE ATNA.
Audit records contain personal data and therefore need their own access controls, retention policy and — because they are the evidence in any investigation — tamper resistance. Append-only storage, or write-once media, or at minimum a separate system with separate credentials.
Detection content that is worth building:
- Access to records of patients with no care relationship
- Access to a VIP or an employee record
- Volume anomalies — a user reading many more records than their peers
- Access outside normal hours or from unusual locations
- Repeated break-glass invocations by the same user
- Failed authorisation clusters
- Large or unusual bulk export requests
A SIEM without health-specific detection content produces alerts about failed logins and misses the clinician browsing their ex-partner's record.
Zero trust
NIST SP 800-207. The operative assumptions: no network location is trusted, every request is authenticated and authorised, and access is granted per session with least privilege.
For health exchange this is simply accurate — the participants are separate organisations. Practical implications: workload identity for services, mutual TLS, per-request policy evaluation at the API layer, and micro-segmentation so that a compromised clinic workstation reaches one facility's systems rather than the national exchange.
Do not attempt it as a single programme. Sequence it: start with strong identity and per-request authorisation at the API gateway, then segment the network, then extend workload identity inward.
Resilience
Security failures and availability failures share a response plan.
Backup. The 3-2-1 rule — three copies, two media, one off-site — plus one immutable or air-gapped copy, because modern ransomware targets backups first. Backups of clinical systems must include the configuration and terminology needed to make the data interpretable, not only the database.
Test restores. On a schedule, to a clean environment, measured. An untested backup is an assumption.
RTO and RPO as clinical parameters. How long can a hospital run without its EMR, and how much data can be lost? These are answered by clinicians, not by IT. A four-hour RTO for an emergency department and a 24-hour RTO for a reporting system are both defensible; a single organisational number is not.
Degraded mode. What happens when systems are down: printed summaries available offline, paper forms, a documented reconciliation process for re-entering data afterwards. Hospitals that survive ransomware well are the ones that had rehearsed running without systems.
Rehearse the incident plan, including the communication path to clinical leadership and to the data protection authority, and including the decision about whether to disconnect from the national exchange — which needs to be a pre-agreed criterion rather than a judgement made under pressure.
Standards and frameworks
| Framework | Use |
|---|---|
| ISO/IEC 27001 | Certifiable management system; often a procurement requirement |
| ISO 27799 | Health-specific application of ISO 27002 controls |
| NIST Cybersecurity Framework | Organising and communicating security posture: govern, identify, protect, detect, respond, recover |
| NIST SP 800-207 | Zero trust architecture |
| CIS Controls | A prioritised, concrete starting list — the most useful framework for a team with limited capacity |
| OWASP ASVS / Top 10 | Application security requirements and common flaws |
| OWASP API Security Top 10 | Directly relevant to FHIR and registry APIs |
For a small team, CIS Controls Implementation Group 1 plus the OWASP API Top 10 delivers more risk reduction per unit of effort than beginning an ISO 27001 certification.
References
- NIST Cybersecurity Framework — https://www.nist.gov/cyberframework
- NIST SP 800-207 Zero Trust Architecture — https://csrc.nist.gov/pubs/sp/800/207/final
- CIS Controls — https://www.cisecurity.org/controls
- OWASP Top 10 — https://owasp.org/www-project-top-ten/
- OWASP API Security Top 10 — https://owasp.org/www-project-api-security/
- OWASP ASVS — https://owasp.org/www-project-application-security-verification-standard/
- FHIR security — https://hl7.org/fhir/security.html
- IHE ATNA — https://www.ihe.net/